Back

Computational Biology and Chemistry

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Computational Biology and Chemistry's content profile, based on 28 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Glycine molecule radical: Predicted properties and dipeptide formation

Synak, J.; Blazewicz, J.

2026-07-10 bioinformatics 10.64898/2026.07.07.736934 medRxiv
Top 0.1%
6.4%
Show abstract

Numerous advances in quantum and computational chemistry over the last decades, well as the development of computer science, allowed utilisation of more precise and complex models, which can be now applied to much bigger systems than in the past. The authors used Gaussian, coupled with theoretical methods, to predict a new way of peptide bond formation, which could have taken place in prebiotic conditions. To better tackle this difficult task, the properties of substrates (glycine-derived radicals) were extensively analysed, using the aforementioned tool - Gaussian, paired with taking resonance and hybridisation into account, to better understand the stereochemistry and the very nature of processes taking place. The result is a series of reactions, which without any sophisticated catalysts and with relatively low energy thresholds ({inverted exclamation}20 kcal/mol) can lead to formation of dipeptides (and further, oligopeptides). The authors also hope, the other predicted properties of the investigated molecules can be of use to any researcher, who would like to utilise them in their experiments. Author summaryOur goal was to investigate a way first peptide bonds in prebiotic conditions could have been formed. This is an extremely important step in research into the beginning of life on Earth. We found a very promising series of reactions, which uses atomic hydrogen as its only catalyst and confirmed our expectations with theoretical calculations, using Gaussian. There are two radicals derived from glycine, which perform major roles in the process, so we investigated their properties with Gaussian and verified that the results are in agreement with our own theoretical considerations. This involved checking for possible geometric isomers and conformers and creating models which could explain their properties. We are well aware that such calculations have limitations and there is no model, which is 100% accurate, so our results should be further confirmed by empirical data in the future. However, we still to be as thorough as possible in how we approached the subject.

2
AptViralDB: A Repository of Experimentally Validated Antiviral Aptamers

Bajiya, N.; Singh, S.; Gahlot, P. S.; Raghava, G. P. S.

2026-07-11 bioinformatics 10.64898/2026.07.08.737144 medRxiv
Top 0.1%
3.6%
Show abstract

In an era of increasing drug resistance, exploring alternative molecules is crucial for the efficient management and treatment of viral diseases. Nucleic acid aptamers have emerged as highly promising candidates due to their exceptional target specificity, low immunogenicity, and versatile mechanisms for viral blocking. This manuscript describes AptViralDB, a manually curated database providing comprehensive information on experimentally validated antiviral aptamers. It contains 1,768 entries of antiviral aptamers against 40 viral species and 104 molecular targets, compiled from literature and existing databases. Each entry provides detailed annotations, including sequence, aptamer type, target, chemical modifications, binding affinity, antiviral activity, stability, and cytotoxicity. We also provide predicted secondary structures and their corresponding minimum free energy (MFE) values. Additionally, a knowledge graph created using ArcadeDB/openCypher enables users to seamlessly explore connections among aptamers, viruses, molecular targets, and biological activities. Finally, the platform offers advanced search and browsing tools, BLAST-based sequence similarity searches, GC-content analysis, downloadable datasets, and REST API access to support computational applications. (https://webs.iiitd.edu.in/raghava/aptviraldb/).

3
Revisiting Logistic Regression for High-Dimensional Gene Expression Data

Souza, R. d. O.; Rodrigues, W. F.; Couto, B.; Dos Santos, M. A.

2026-07-24 bioinformatics 10.64898/2026.07.20.739668 medRxiv
Top 0.2%
2.6%
Show abstract

Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.

4
A Curvature Guided Composite Kernel Framework for Differential Gene Selection in Cancer Transcriptomics

Gupta, M.; Sarkar, A. P.

2026-07-28 bioinformatics 10.64898/2026.07.23.740443 medRxiv
Top 0.2%
2.6%
Show abstract

Identification of differentially expressed genes is a crucial step for downstream tasks on gene data such as biomarker discovery, drug target identification. Traditional Methods assume negative binomial distribution on RNA-sequence data and models the DEGs using either generalized linear models or by estimating dispersion and assumption of mean-variance rate. The proposed method uses axiomatic approach by using quantum mechanics principles to project transcript data onto a Hilbert space using a composite kernel. Using the curvature generated by the transcripts on the latent manifold within the Hilbert space, a gravitational search inspired mechanism is used to identify the optimal number of differentially expressed genes by minimizing a representational loss function, and a reduced gene feature space is constructed as the potential differentially expressed genes. The proposed method has been compared with existing empirical methods for validation using proper statistical and biological benchmark analysis.

5
Global protein expression profiling in stem cell factor stimulated human Acute megakaryoblastic leukemia cells identifies CFL1, GSN and CCT8 as prognostic biomarkers for Acute Myeloid Leukemia.

Ravi, A. K.; Gopan, G.; Arumugam, S.; Sethumadhavan, A.; Mani, M.

2026-08-26 cancer biology 10.64898/2026.08.24.746695 medRxiv
Top 0.3%
2.0%
Show abstract

Abstract Background: The stem cell factor receptor or c-Kit is a type III receptor tyrosine kinase, activated by its ligand Stem cell factor (SCF). Up on activation, c-kit induces signaling pathways that regulates blood cell proliferation, survival, differentiation, and migration. Several studies reported that c-Kit/SCF signaling, contributes to the development and progression of acute myeloid leukemia (AML) in patients. However, the downstream proteins regulated by c-kit activation and their clinical significance in AML remain poorly explored. Methods: Human Acute megakaryoblastic leukemia (Mo7e) cells, were-stimulated with SCF and global protein expression were profiled using two-dimensional gel electrophoresis coupled with MALDI-TOF and LC-MS/MS. Differentially expressed proteins were functionally characterized and validated using patient data from the TCGA-LAML and matched normal data from GTEx, GEO datasets, and quantitative RT-PCR. Their diagnostic and prognostic significance was assessed using ROC, Cox regression, LASSO, Kaplan Meier survival analyses, and a prognostic nomogram model. Results: Proteomic profiling identified 14 differentially expressed proteins in SCF-stimulated Mo7e cells, which are predicted to involved in cytoskeletal organization, protein folding, metabolism, vesicular trafficking, and translational regulation. Transcriptomic analysis of the TCGA-LAML cohort revealed significant dysregulation of CFL1, CCT8, HSP90B1, MDH2, EIF5A, GSN, and TPI1. Integrated ROC, Cox regression, and LASSO analyses identified CFL1, CCT8, and GSN as the most robust prognostic biomarkers associated with poor overall survival in LAML patients. Their expression patterns were validated in independent GEO datasets and by qRT-PCR in SCF stimulated Mo7e cells. Finally, a three-gene nomogram model was developed and validated to predict the overall survival probability of AML patients at 1-, 3-, and 5-year time points. Conclusions: This study identifies CFL1, CCT8, and GSN as key downstream effectors of c-Kit signaling as prognostic biomarkers for AML. These findings provide mechanistic insights into c-Kit-driven leukemogenesis and establish a clinically relevant three-gene signature for AML risk stratification and potential therapeutic targeting.

6
AptCancerDB: A Curated Knowledgebase and Translational Discovery Platform for Anticancer Aptamers

Bajiya, N.; Singh, S.; Raghava, G. P. S.

2026-07-09 cancer biology 10.64898/2026.07.02.735999 medRxiv
Top 0.3%
2.0%
Show abstract

Aptamers are emerging as important molecular recognition ligands in oncology, playing significant roles in cancer diagnostics, targeted therapies, drug delivery systems, and molecular imaging. Numerous aptamers have advanced to clinical trials, indicating their potential for real-world applications; however, existing databases fail to capture that. To bridge this critical gap, we developed AptCancerDB (https://webs.iiitd.edu.in/raghava/aptcancerdb/), a comprehensive, manually curated database of experimentally verified anticancer aptamers. The current release contains 1,941 entries collected from studies published between 2000 and 2025, covering 29 cancer types, approximately 200 cancer cell lines, and direct links to 22 clinical trials. Each entry is annotated with sequence information, target details, cancer type, cell line, SELEX methodology, affinity determination data, chemical modifications, and biological activities. The dataset is dominated by 82.7% ssDNA, reflecting its superior stability and ease of synthesis, while only 16.6% is ssRNA and appears primarily in studies targeting complex intracellular or protein-protein interactions. To facilitate structural analysis, predicted secondary structures, dot-bracket notations, specific structural elements, and minimum free energy values were also included. AptCancerDB integrates a MySQL backend with an ArcadeDB/OpenCypher-based Knowledge Graph, enabling exploration of relationships among aptamers, targets, cancer types, cell lines, and functional applications. The platform provides advanced search and browsing facilities, BLASTn-based similarity searching, and GC Calculator. Built on a modern, responsive frontend (React/TypeScript/Tailwind CSS), the platform includes a REST API for data retrieval. By integrating fragmented experimental data into a unified cancer-focused resource, AptCancerDB serves as a valuable resource for comparative analysis, aptamer discovery, and the development of next-generation aptamer-based diagnostics and therapeutics. HighlightsO_LICurated knowledge base of experimentally validated anticancer aptamers. C_LIO_LIAptCancerDB contain therapeutic, tumor-homing and cell-penetrating aptamers. C_LIO_LISummarizes clinical progress and translational trends in anticancer aptamer research. C_LIO_LISupports rational aptamer design using molecular, functional, and clinical annotations C_LIO_LIDisease-focused resource for cancer diagnosis, therapy, and drug delivery C_LI TeaserAptCancerDB maintains experimentally validated anticancer aptamers relevant to diagnosis, drug delivery, and therapy.

7
The Origin and Evolution of Protein Synthesis: A Co-Adaptation Flexible-Rigid Docking Model Based on First-Principles Reasoning

Zhao, D.; Yang, Y.; Sun, J.; Zhang, J.; Duan, H.; Tan, Y.; Liu, l.

2026-06-11 evolutionary biology 10.64898/2026.06.11.730790 medRxiv
Top 0.3%
1.9%
Show abstract

Although the "RNA world" hypothesis suggests that RNA played a crucial role in the origin of life [7], the functional framework of RNA in prebiotic protein synthesis and the mechanisms of genetic code formation during the prebiotic period remain poorly understood. Here, using the prebiotic "primordial soup" as a model, we reconstructed the detailed steps that would yield a protein with a stable ordered amino-acid sequence in the "primordial soup" at the prebiotic period. In the "primordial soup", a large number of medium- to large-sized biomolecule-like substances--such as RNA-like and protein-like molecules of various sizes and shapes, as well as related polymers like amino-acid-RNA-like etc.--did generate and accumulate. Moreover, protein-like and RNA-like molecules formed even more intricate complexes. These complexes bound free mRNA-like molecules through complementary base pairing. Subsequently, with an extremely low probability, two adjacent amino-acid-RNA-like molecules became bound to this free mRNA-like molecule, and their amino acids underwent a condensation reaction by the complexes, producing peptides and eventually proteins or polypeptides. This free mRNA-like molecule exhibits a certain flexible structure, whereas the super-large complexes formed by protein-like and RNA-like molecules (which possess certain activities) and the amino-acid-RNA molecules exhibit relatively rigid structures. Long-term evolution and mutual selection led to the emergence of proteins with stable amino acid sequences and moderate catalytic activity. In this way, the nucleotide information embedded in such mRNA-like molecules indirectly express through protein synthesis--a process we term the "A Co-Adaptation Flexible-Rigid Docking Model", where flexible mRNA-like molecules dock onto rigid complexes to enable ordered peptide formation. Finally, we show how trinucleotide codons emerge naturally from the flexible-rigid docking constraints.

8
EnzyKAN: Protein Language Model Embeddings and Kolmogorov-Arnold Network Variants for Enzyme Commission Classification with a Proposed Electron-Transfer Physics Feature Framework

R, S.; Reddy, B. R. R.

2026-06-29 bioinformatics 10.64898/2026.06.23.734004 medRxiv
Top 0.3%
1.9%
Show abstract

MotivationComputational enzyme classification has previously utilised sequence homology features and protein language model embeddings. The Kolmogorov-Arnold Network (KAN) paradigm, which uses learnable edge functions rather than fixed ones, has shown promising results in biological sequence tasks. ResultsA fully reproducible investigation of KAN variants for seven-class EC classification on up to 9,516 labelled sequences from the CLEAN benchmark [1] (9,386 for language model experiments). In the sequence only settings, fixed basis KAN variants outperformed an MLP baseline moderately (macro F1 = 0.17-0.29). Utilisation of ESM-2 650M embeddings [2] greatly improved results via 5-fold cross-validation: MLP macro F1 = 0.750 {+/-} 0.009, accuracy = 0.823 {+/-} 0.009; learnable SineKAN macro F1 = 0.716 {+/-} 0.023, accuracy = 0.788 {+/-} 0.019. MLP performed comparably but did not exceed conventional baselines. As an aside, we introduce but do not investigate an approach to EC oxidoreductase sub-classification through the use of a Marcus theory-based electron transfer feature framework. AvailabilityCode and result files are available at https://github.com/sanjuz-cas/ENZYKAN.

9
Plasma proteome screen of individuals with elevated blood counts to identify diagnostic signatures and biomarkers inherent to patients with myeloproliferative neoplasms

Zhong, X.; Lundahl, I.; Rosell, A.; Chaireti, R.; Ungerstedt, J.

2026-08-04 cancer biology 10.64898/2026.08.03.742450 medRxiv
Top 0.3%
1.8%
Show abstract

BackgroundThe Philadelphia negative myeloproliferative neoplasms (MPN), including essential thrombocythemia (ET), polycythemia vera (PV) and primary myelofibrosis (PMF), are characterized by myeloid cell proliferation, thrombosis and inflammation. Suspicion of MPN arises from increased blood count in one or more lineages; however, knowledge on the MPN plasma proteome including biomarkers measurable in blood, are lacking. Comparing the plasma proteome of MPN patients to subjects with elevated blood counts but without MPN diagnosis, may provide an increased understanding of the MPN disease biology as well as diagnostic biomarkers measurable in blood. Patients and methodsWe performed plasma proteome profiling in 87 patients referred to the Department of Hematology due to elevated blood counts. Of these, 55 were diagnosed with MPN and 32 did not fulfill MPN diagnostic criteria and thus constituted the non-MPN control group. ResultsThe frequency of thrombosis was equal between the groups. We found 189 differentially expressed proteins between MPN and non-MPN, enriched for Hemostasis and Platelet activation proteins. Using Lasso multiple regression, we identified SORT1, GP1BA, PSPN, MMP1 and BAG6 separating MPN from non-MPN individuals, and TFRC and SEMA7A specific for MPN subtype PV. Interestingly, TFRC alone had a diagnostic accuracy for identifying PV of 89.3%, and when combined with serum erythropoietin it increased to 99.3%. Only two proteins, IL-6 and GH1, were increased in JAK2 mutant MPN compared to JAK2 wildtype MPN. The same trend was seen for JAK2 mutant ET compared to JAK2 wildtype ET, indicating that IL-6 is induced by JAK STAT activation. However, IL-6 levels did not differ between MPN and non-MPN patients. Discussion/conclusionIn conclusion, hemostasis and platelet activation are inherent to MPN disease whereas little difference was found in proinflammatory cytokines between MPN and non-MPN groups. We demonstrate novel potential blood biomarkers for MPN and MPN subtypes, in particular TFRC for identifying PV patients.

10
Functional Characterization of Transcriptome-Wide Isoform Switching in Hürthle Cell Carcinoma (HCC)

Butt, R. S.; Amir, A.; Paracha, R. Z.

2026-07-27 bioinformatics 10.64898/2026.07.23.740299 medRxiv
Top 0.3%
1.8%
Show abstract

Hurthle cell carcinoma (HCC) is an aggressive form of thyroid cancer. While mitochondrial DNA mutations and chromosomal losses have been identified in HCC, isoform switching, and its functional consequences remain uncharacterized. This study reanalyzed NCBI GEO dataset GSE228870 (n = 32), using Salmon and IsoformSwitchAnalyzeR() to identify isoform switching. The analysis resulted in 371 switches across 335 genes showing functional consequences including loss of protein domains, shorter open reading frames (ORFs), loss of signal peptides and novel sub-cellular localizations. Most significant isoform switches (q-value < 0.05, |dIF| > 0.1) were observed in LAMA2, LSP1, MAD2L2, FBLN2 and CXCL12, implicating extracellular matrix dysregulation, DNA damage response, immune signaling and cytoskeleton regulation. These genes are expressed in normal thyroid (median TPM 20.69, 11.66, 14.79, 134.1 & 80.76). However, specific isoforms of LAMA2 and MAD2L2 are not expressed in normal thyroid, explaining tumor-specific expression in HCC. Alternative transcription termination site (ATTS) gain was significant, suggesting altered 3 end in HCC transcripts. TCGA SpliceSeq showed LSP1, FBLN2 and CXCL12 undergo alternative promoter (LSP1 exon1 PSI=94.5%, FBLN2 exon2 PSI=99.0%) and alternative termination (CXCL12 exon3.3 PSI=53.9%) in thyroid cancer, suggesting ATTS and alternative transcription start site (ATSS) as shared splicing dysregulation mechanisms. This is the first systematic characterization of isoform-level dysregulation in HCC.

11
MEGAHIT k-mer range tuning trades computational efficiency for improved recovery of functional genes across cave sediment and wastewater metagenomes

Carunta, A.; Banciu, H. L.; Mizeranschi, A. E.

2026-07-17 bioinformatics 10.64898/2026.07.17.739120 medRxiv
Top 0.4%
1.7%
Show abstract

Shotgun metagenomics is a powerful approach for profiling complex microbial ecosystems and discovering functional genes, including antimicrobial resistance genes (ARGs) and biosynthetic gene clusters (BGCs). De novo assembly with tools such as MEGAHIT commonly uses multiple k-mer lengths, but the effect of reduced k-mer sets on functional gene recovery has received limited attention. Here, we quantify the trade-off between assembly speed and functional-gene recovery using 17 cave sediment metagenomes and 10 wastewater metagenomes assembled under 19 MEGAHIT k-mer scenarios. In cave metagenomes, finer-grained k-mer ranges recovered more BGCs and, in several pairwise comparisons, more ARGs, but required longer runtimes. In wastewater metagenomes, finer-grained settings most clearly affected ARG recovery, whereas BGC counts did not differ significantly after Friedman testing. These results indicate that reduced k-mer sets can lower computational cost but may miss biologically relevant functional signal, depending on the dataset and downstream target. The study provides a quantitative basis for selecting MEGAHIT k-mer parameters according to whether computational efficiency or functional gene discovery is the primary aim.

12
A Simple Method to Distinguish Active and Inactive Aptamers by Analyzing the Ruggedness of the Aptamer Free Energy Landscape

Subramanian, G.; Thiel, W.; Singh, R.

2026-08-29 bioinformatics 10.64898/2026.08.26.747184 medRxiv
Top 0.4%
1.7%
Show abstract

Aptamers are structured nucleic acid ligands capable of high affinity, high specificity molecular recognition generated using variations of the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) process. However, SELEX often produces sequences that enrich yet may lack binding efficacy. We propose a measure called the Ruggedness Composite Index (RCI) along with a method for computing it, that can be used to distinguish binding-competent ('active') aptamers from weak or non-binding ('inactive') aptamers. Given a set of aptamers, RCI incorporates information on their fragmentation (landscape partitioning), basin entropy (metastable state distribution), cumulative density irregularity (non-uniform occupancy), and structural energy correlation length (structure-energy coupling scale). We test whether secondary-structure folding energy landscape topology distinguishes active from inactive aptamers using a multiscale level set framework across six datasets. Active aptamers show lower RCI values and occupy smoother, funnel-like conformational spaces, while inactive aptamers show higher RCI values, reflecting fragmented, high-entropy landscapes. By contrast, classical thermodynamic features, such as minimum free energy, show limited discrimination between active and inactive aptamers. In all datasets, sequences that exhibit enrichment which is not monotonic but lack specificity exhibit elevated ruggedness, indicating landscape topology can predict non-specific enrichment. These results indicate that folding landscape organization can be used as a predictor of aptamer activity and establish RCI as a simple, mechanistically interpretable measure for improving candidate prioritization, especially in therapeutic aptamer discovery.

13
Nickel-Driven Dynamics of Urease in Sporosarcina pasteurii: Integrated Computational and Experimental Insights

Al-Thawadi, S. M.

2026-06-19 bioinformatics 10.64898/2026.06.15.732323 medRxiv
Top 0.4%
1.5%
Show abstract

Urease is a nickel-dependent enzyme that plays an important role in urea hydrolysis and in a process named as microbial-induced calcium carbonate precipitation (MICP), which is widely used in sustainable environmental biotechnology. Despite its ecological importance, urease powers Biogrout (biocementation), a promising green technology for soil stabilization and infrastructure repair. Yet, the relationship between nickel availability, enzyme activation, and bacterial fitness remains poorly understood. In this study, we reveal a striking dual effect of nickel on Sporosarcina pasteurii: while high Ni{superscript 2} concentrations strongly inhibit growth (IC {approx} 637.7 {micro}M), they simultaneously boost specific urease activity up to six-fold. This uncoupling between biomass and enzymatic efficiency highlights a previously overlooked adaptive strategy under metal stress. Using structural bioinformatics and molecular docking, we show that Ure1--the catalytic subunit--exhibits the strongest nickel affinity (-4.3 kcal{middle dot}mol-{superscript 1}), supported by highly conserved active-site residues, whereas accessory proteins UreE and UreG display moderate and weak binding, consistent with their roles in metal delivery and GTP-dependent maturation. In addition, microscopic observations confirmed that calcium carbonate precipitation was most pronounced at intermediate nickel concentrations (approximately 400-1000 {micro}M), whereas higher concentrations ([&ge;]1000-1300 {micro}M) led to reduced mineral formation due to loss viable cells. Taken together, these results indicates that nickel availability controls both urease activation and bacterial fitness, and that an optimal balance is required to maximize biomenerilization efficiency in environmental applications, particularly in biocementation technology. ImportanceUrease-driven biomineralization is widely used in sustainable technologies such as soil stabilization and self-healing concrete. However, optimizing these systems requires a clear understanding of how environmental factors influence enzyme performance. This study shows that nickel, an essential cofactor for urease, plays a dual role by enhancing enzymatic activity while inhibiting bacterial growth at high concentrations. By integrating experimental data with computational analysis, we demonstrate that efficient biomineralization depends on maintaining nickel within an optimal range that balances enzyme activation and microbial viability. These findings provide practical guidance for improving biocementation processes and highlight nickel as a key regulator of urease-based environmental biotechnology applications.

14
LsGCRPred: lncRNA-SNP regulated gene expression-based breast and ovarian cancer risk prediction model

Das, T.; Das, G.; Ghosh, B.; GHOSH, Z.

2026-07-07 cancer biology 10.64898/2026.07.07.736696 medRxiv
Top 0.4%
1.5%
Show abstract

Long non-coding RNAs (lncRNAs) and single nucleotide polymorphisms (SNPs) within them play crucial role in cancer susceptibility and disease outcomes. Breast and ovarian cancers, characterized by genetic heterogeneity, present significant challenges for precise diagnosis and treatment. Despite recent advancements in personalized medicine, inclusion of lncRNA-SNP (LSNP) markers into cancer risk detection panels remains limited. In this work, we put forward lncRNA-SNP regulated gene expression-based breast and ovarian cancer risk prediction model LsGCRPred (LSNP-Gene Interaction Based Cancer Risk Prediction Model). Notably, our approach accounts for the tissue-specificity of lncRNAs as well the benefit for individuals with predisposing conditions. Additionally, pathway analysis revealed the involvement of the LSNP interacting genes in key cancer regulating pathways. TaqMan genotyping and qPCR were performed to confirm the presence of selected LSNPs in ovarian and breast cancer cell lines along with the significant expression of the lncRNA and associated gene transcripts. These findings highlight previously overlooked genetic variants within lncRNA loci and their regulatory impact on disease outcomes, providing insights into personalized cancer diagnosis and treatment strategies. The tool LsGCRPred can be accessed as a standalone version on GitHub. Github Link: https://github.com/zglabDIB/LsGCRPred

15
Identification of Altered Potassium Channels for Drug Repurposing in Long COVID Patients

George, J. P.; Gaikwad, K. B.; Sharma, J.

2026-06-19 bioinformatics 10.64898/2026.06.18.733062 medRxiv
Top 0.5%
1.4%
Show abstract

Long COVID (LC) is a complex condition characterized by persistent, chronic multisystem manifestations, with a significant proportion of patients exhibiting neurological symptoms. Human ion channels (HICs), particularly potassium channels, are abundantly expressed in the nervous system and linked to key metabolic processes, making them potential candidates for understanding LC pathophysiology and drug repurposing. Meta-analysis of RNA-Seq datasets from COVID-19 recovered and LC patients was performed to identify altered HICs in LC. Differential gene expression analysis, functional enrichment analysis, and weighted gene co-expression network analysis (WGCNA) were performed to uncover key genes, pathways, and co-expression modules consisting of HICs, lipid metabolism-, and immune signaling-related genes. Drug-gene interaction analysis was performed to identify approved drugs targeting potential HICs. A total of 715 dysregulated genes, including eighteen HICs were identified, among which seven were potassium channels. Three significant modules containing HICs, lipid metabolism-, and immune signaling-related genes were identified and found to be associated with antigen processing and presentation, complement and coagulation cascades, and cytokine-related pathways. Approved drugs targeting KCNA6, KCNJ10, KCNN3, and KCNH4 were identified. With further experimental validation, these dysregulated potassium channels, supported by their co-expression networks and pathway associations, may act as potential candidates for drug repurposing in LC patients.

16
Immunoinformatics-Guided Design and In Silico Evaluation of a Multi-Epitope Vaccine Against Influenza A H10N5 and H3N2 Strains Based on Hemagglutinin and Neuraminidase Proteins

Shabbir, M. Z.; Kumar, P.; Rehman, M. A. U.; Kumar, J.; Urooj, U.; Batool, S. I.; Sourav, C.; Ghazanfar, R.; Nagari, Z.; Hameed, D.; Wahid, A.; Atique, A.; Siddique, M. D.

2026-07-08 bioinformatics 10.64898/2026.07.03.736294 medRxiv
Top 0.5%
1.4%
Show abstract

Influenza A viruses H3N2 and H10N5 represent, respectively, a persistently dominant seasonal pathogen and a newly documented zoonotic threat with the latter strain variants responsible for the first confirmed human fatality in January 2024, yet no vaccine platform currently addresses co-protection against both subtypes within a unified immunogen. We report here the immunoinformatics based vaccine design and multi-layered computational validation of a 419-amino-acid multi-epitope subunit vaccine construct targeting conserved hemagglutinin (HA) and neuraminidase (NA) antigens identified through multiple sequence alignment of the avian H10N5 (A/swine/Hubei/10/2008) and H3N2 human reference strain sequences to identify viral agents undergoing mammalian adaptations. Linear B-cell, cytotoxic T lymphocyte (CTL), and helper T lymphocyte (HTL) epitopes were predicted using ABCpred, BCEpred, BepiPred 2.0, NetMHCpan 2.1, and NetMHCpan 4.0, then filtered through VaxiJen 3.0, AllerTOP v2.1, and ToxinPred to retain only antigenic, non-allergenic, non-toxic candidates. The final construct, incorporating an avian {beta}-defensin N-terminal adjuvant with GPGPG, AAY, and EAAAK linkers, exhibited a molecular weight of 43.9 kDa, instability index of 31.15, and SOLPro solubility probability of 0.763. Tertiary structure modeling via I-TASSER and GalaxyRefine achieved 84.4% Ramachandran-favored residues. Molecular docking against TLR3 and TLR7 yielded binding free energies of -16.1 and -16.8 kcal/mol with picomolar dissociation constants. Molecular dynamics simulations confirmed complex stability over extended trajectories. Furthermore, codon optimization produced a Codon Adaptation Index of 1.0 for E. coli K12 expression. In silico immune simulation demonstrated robust activation of humoral and cellular immunity including elevated IgG1, IgM, IFN-{gamma}, IL-2, rapid NK cell expansion, and broad B-cell clonal diversity. These findings establish a computationally validated candidate capable of providing protection against influenza in multiple host organisms, warranting experimental advancement.

17
Computational and Structure-Guided E-Pharmacophore-Based Virtual Screening for the Identification of Novel NEK2 Kinase Inhibitors as Potential Anticancer Agents

Rehman, H. M. M.; Latif, A.; Hammad, H. M.; Sajjad, M.

2026-08-06 bioinformatics 10.64898/2026.07.31.742111 medRxiv
Top 0.5%
1.4%
Show abstract

Cancer is a serious public health problem, and is becoming more common, with a projected increase in deaths and more than 25 million new cases in coming decades. A number of molecular mechanisms are involved in the tumoral process, with one of them, never in mitosis A-related kinase 2 (NEK2), a serine/threonine protein kinase, being a frequent target of amplification in various malignancies that is responsible for chromosomal instability, aneuploidy and activation of several oncogenic pathways. Available kinase inhibitors are not yet optimized in terms of their pharmacokinetic properties for clinical use, and current therapies, such as chemotherapeutic agents or immunotherapies are often limited by their resistance. In silico methods represent an effective tool to search for novel potent inhibitors, before testing in animals, with time constraints and limited resources. To find new inhibitors of NEK2, we used E pharmacophore-based modeling and structure based virtual screening in this study. NEK2 was chosen as the target for therapeutic intervention and an energy optimized pharmacophore model was employed to screen the Enamine REAL library of millions of compounds. Pharmacodynamic and Pharmacokinetic properties of the Top hits were tested using ADMET profiling. These were further screened using molecular docking (standard precision and extra precision) and virtual screening to obtain three lead compounds 1, 2, and 3 which have docking score of -7.414, -8.037 and -7.562 respectively. MM-GBSA calculations were used to estimate the binding free energies for these complexes, which were determined to be -54.92, -54.18 and -49.23 kcal/mol. Lastly, 100 ns molecular dynamics simulations have been run to evaluate complex stability in dynamic situations. The overall results of the MD showed the overall stability of the NEK2-ligand complexes, and thus these three compounds are promising NEK2 inhibitor candidates and could be further validated in vitro and in vivo for clinical application. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=133 SRC="FIGDIR/small/742111v1_ufig1.gif" ALT="Figure 1"> View larger version (70K): org.highwire.dtl.DTLVardef@14e9cfeorg.highwire.dtl.DTLVardef@24ec78org.highwire.dtl.DTLVardef@20be75org.highwire.dtl.DTLVardef@1b80c02_HPS_FORMAT_FIGEXP M_FIG C_FIG

18
Trinucleotide Distribution, Symmetry Elements and Formulation of Mirror Symmetry Index for G4 Motifs

Arya, A.; Datta, B.

2026-07-05 bioinformatics 10.64898/2026.07.05.736592 medRxiv
Top 0.5%
1.3%
Show abstract

Symmetry elements in nucleic acids are most strongly correlated with sites of biological function; however, their relevance to non-canonical structures remains underexplored. In this study, we demonstrate the presence and significance of trinucleotide symmetry elements within G-quadruplex (G4) motifs. Our central hypothesis is that the intra-strand mirror symmetry of trinucleotides has been evolutionarily selected to facilitate G4 formation builds on the established sequence-structure association of G-quadruplexes and the natural symmetry law governing nucleotide insertion during genome evolution. Using a conserved G4 motif in the first exon of the MTOR gene as a model, we showed remarkable trinucleotide symmetry preservation across primates and broader mammals, with functional G4 regions displaying locally elevated symmetry relative to the codon-biased exonic background. Analysis of experimentally validated oncogenic G4s, including c-MYC, BCL2, VEGF, and KRAS, revealed that mirror and reverse complement symmetries converge around biologically important G4s. To quantify this feature, we formulated two complementary descriptors: the mirror symmetry index (MSI) and its non-palindromic variant (nMSI). Across 14 oncogene-promoter wild-type G4s, the majority scored MSI [&ge;] 0.80 (mean 0.884), with only the loop-rich ATG7, BCR, and MDM2 motifs falling below this value, and the KRAS promoter G4 reached individual significance against its mononucleotide-preserving null distribution (p = 0.042). Most decisively, each wild-type G4 scored higher on MSI than its experimentally confirmed G4-abolished mutant in 12 of 14 paired comparisons (sign test, p = 0.0065; mean {Delta}MSI = +0.089, mean {Delta}nMSI = +0.192); the two reversals (BCL2 and HIF-1) are attributable to scrambled mutant controls that introduce more balanced trinucleotide compositions rather than to failure of the index. The directional trend was reproduced across three independently published datasets, with nMSI [&ge;] 0.50 separating G4-forming from non-G4 sequences at 77.8% sensitivity and 100% specificity, although the collective per-sequence signal from mononucleotide-preserving shuffles remained a non-significant trend (Stouffer combined Z = 1.197, p = 0.116). This first report of trinucleotide symmetry in G4 motifs posits that coordinated nucleotide insertion and quadruplet maintenance act as an evolutionary forcing mechanism that pre-organizes single strands for G4 folding.

19
Evaluation of Trypanosoma brucei Phosphofructokinase Allosteric Inhibition: An In-Silico Study

Gumbis, G.; Houston, D. R.

2026-06-20 bioinformatics 10.64898/2026.06.16.732740 medRxiv
Top 0.6%
1.3%
Show abstract

Human African trypanosomiasis, caused by a protozoan parasite Trypanosoma brucei, is a neglected tropical disease for which well-tolerated, conveniently administered, and highly efficacious medicines are still missing. Previously, T. brucei Phosphofructokinase was targeted by small-molecule inhibitor development efforts. This approach has shown promise both in vitro and in vivo. In this study, we have used these wet-lab results, evaluated the compounds already characterised by Molecular Dynamics simulations, found relationships between in silico and wet-lab data and used these observations to evaluate compounds that we selected through several different approaches of virtual screens. We observed that inhibitor-ATP interactions are highly predictive of the inhibitory activity. Several compounds selected through virtual screens have outperformed previously characterised compounds.

20
Artificial Intelligence Model: Optimizing Cancer Risk Level Predictions Using Machine learning and deep learning approaches

Abd Aziz, A. B.; Arabiat, A.; Abu Owida, H.; Abuowaida, S.; Alshdaifa, N.; A. Mashagba, H.

2026-08-25 cancer biology 10.64898/2026.08.20.745910 medRxiv
Top 0.6%
1.2%
Show abstract

This study emphasizes the potential of computational techniques in cancer risk assessment, lighting opportunities for specific and data-driven healthcare solutions. This study examines the use of artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches to improve cancer risk assessment using a Kaggle dataset. The study uses Java-based ML software to create and evaluate multiple predictive models, taking advantage of its powerful libraries and frameworks for processing and analyzing cancer risk indicators. This work analyzes model performance using 10-fold cross-validation, resulting in reliable generalization and accuracy estimates. Several classification techniques, such as Random Forest (RF) logistic regression (LR), decision trees (DT), Naive Bayes (NB), and Multi-layer perceptron (MLP), are used to assess their efficacy in predicting risk levels for various cancer types. To measure classification effectiveness, key performance metrics such as accuracy, precision, recall, and F1 score are produced, in addition to multi-class confusion matrices. The results show that the RF model is the best classifier for classification, with accuracy of 99.85%, F-measure of 99.80%, precision of 99.80%, and sensitivity of 99.90%. These findings demonstrate the model's ability to effectively estimate cancer risk levels among individuals. of cancer risk estimations, allowing for earlier discovery and more effective medical care.